Back

Journal of the American Medical Informatics Association

Oxford University Press (OUP)

All preprints, ranked by how well they match Journal of the American Medical Informatics Association's content profile, based on 71 papers previously published here. The average preprint has a 0.14% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
Embedded point of care stratified block randomization: demonstration of the Point of Care Randomization (POCR) engine with an electronic health record pragmatic clinical trial

Sarkisian, C.; Ibrahim, K.; Vangala, S.; Villaflores, C. W.; Cheng, E. M.; Turner, W.; Leuchter, R. K.; Machado, A.; Tabar, J.; Verdeflor, J. A.; Purvis, J.; Goncharova, A.; Pletcher, M. J.

2026-01-28 health informatics 10.64898/2026.01.26.26344847 medRxiv
Top 0.1%
70.5%
Show abstract

We describe a new custom feature within our Epic Systems electronic health record (EHR) that automates stratified randomization at the point-of-care or order. As a demonstration use-case, we conducted a randomized trial of a provider-facing alert for short-interval HbA1c orders. Over 3 months the alert dramatically reduced repeat orders. This transportable clinical informatics application transforms health systems ability to conduct pragmatic clinical trials and deliver clinical care within the EHR.

2
Care Plan Generation for Underserved Patients Using Multi-Agent Language Models: Applying Nash Game Theory to Optimize Multiple Objectives

Basu, S.; Baum, A.

2026-02-25 health informatics 10.64898/2026.02.23.26346934 medRxiv
Top 0.1%
64.9%
Show abstract

BackgroundClinicians in care management programs are often in low supply relative to patient demand, especially in US Medicaid programs, and must simultaneously address clinical risk, time efficiency, and patients social needs. Many studies have shown that large language models may assist in their tasks for summarizing patient care, such as in generating care plans; yet these studies also show that different objectives given to agents often conflict and produce problems for safety, efficiency and equity. We tested whether and to what degree using game theoretic approaches (a Nash bargaining framework) can produce care plans that advance multiple objectives across multiple language models, applying data from a real-world Medicaid cohort. MethodsWe conducted two studies in a cohort of 5,148 activated Medicaid care management patients (69.9% female; 45.7% Black or African American; mean age 40.9 years) enrolled in Virginia and Washington. A retrospective evaluation applied five deterministic strategies to the full cohort to characterize multi-objective trade-offs. A pre-registered controlled paired experiment (N = 200) assigned each patient one Nash-orchestrated multi-agent plan and one compute-matched sequential self-critique plan, generated by locally hosted open-source models (DeepSeek-R1 8B; Llama 3.1 8B) with no patient data leaving local infrastructure. Pre-specified outcomes were Safety, Efficiency, Equity, and Composite (mean of the three), each scored 0-1. Reporting follows CONSORT 2010 and STROBE. ResultsNash orchestration produced a Composite score of 0.755 (95% CI 0.751-0.760) versus 0.742 (95% CI 0.739-0.746) for the compute-matched baseline; the paired difference was 0.013 (95% CI 0.008-0.019; p = 6.20 x 10-). Safety and Efficiency paired differences were small-to-moderate in effect size (Cohens d = 0.327 and 0.543, respectively) with confidence intervals excluding zero. The Equity paired difference was 0.000 (95% CI -0.015 to 0.014; p = 0.987). ConclusionsRole-specialized Nash-orchestrated multi-agent language models produced measurably better Safety and Efficiency care plan quality than a compute-matched baseline under data-residency constraints. The null Equity result demonstrates that multi-objective role specialization does not automatically address equity--equity requires explicit design attention beyond composite weighting--with direct implications for responsible AI deployment in Medicaid care management. Author SummaryCare management programs for Medicaid patients need to address multiple goals at once: covering clinical risks, prioritizing the most impactful interventions, and recognizing the social barriers that affect whether patients can follow through on care plans. Prior research shows that automation tools powered by a single AI model tend to optimize for one of these goals at a time, sacrificing the others. We tested whether organizing several specialized AI agents -- each focused on a different goal -- and then combining their recommendations through a mathematical framework called Nash bargaining could produce better overall care plans for a real Medicaid population. We found that this multi-agent approach produced care plans that the AI judge rated as meaningfully safer and more efficient than plans generated by a single AI model using the same total amount of computation. However, the multi-agent approach did not produce plans that were more equitable in addressing patients social needs, suggesting that equity requires more direct attention as a design target rather than emerging from multi-objective combination alone. All AI inference was performed on locally hosted computers, with no patient information sent to outside services, reflecting the privacy requirements of real-world Medicaid care management programs.

3
Identifying Reasons for ACEI/ARB Non-Use in CKD Using Scalable Clinical NLP with Schema-Guided LLM Augmentation

Al-Garadi, M.

2026-02-12 health informatics 10.64898/2026.02.10.26346025 medRxiv
Top 0.1%
60.2%
Show abstract

IMPORTANCEAlthough angiotensin-converting enzyme inhibitors (ACEIs) and angiotensin receptor blockers (ARBs) are recommended for people with chronic kidney disease (CKD), they remain underused. Barriers to adherence, such as adverse effects or patient refusal, are frequently embedded within unstructured clinical narratives and are therefore inaccessible to structured data analytics. Scalable natural language processing (NLP) approaches are needed to identify these barriers and support guideline-concordant care. OBJECTIVETo develop and evaluate an NLP model capable of identifying documented reasons for ACEI/ARB non-use within clinical notes of people with CKD in the Veterans Affairs (VA) healthcare system. DESIGN, SETTING, AND PARTICIPANTSThis retrospective study analyzed electronic health record data from 2005 to 2024 including people aged 18 to 80 years with CKD, defined by an estimated glomerular filtration rate (eGFR) of 20-60 mL/min/1.73 m2 and presence of albuminuria, across multiple VA medical centers. NLP models were trained on 1,025 manually annotated notes and further augmented with 4,600 synthetic examples generated through schema-guided large language model prompting. MAIN OUTCOMES AND MEASURESThe primary outcome was model performance in identifying notes containing at least one documented reason for ACEI/ARB non-use, evaluated using F1-score, precision, and recall. Secondary outcomes included model learning curve analyses and the effect of synthetic data augmentation on classification performance. RESULTSThe most common documented reasons for ACEI/ARB non-use were acute kidney injury (29.6%), increased creatinine (12.4%), cough (11.2%), and hypotension-related symptoms (11.1%). Across modeling approaches, training with synthetic data augmentation improved detection of notes containing reasons for non-use. Performance gains were statistically significant across all models (McNemar test, P < .05), with the random forest model using Nomic embeddings achieving the highest performance (F1 score, 0.79; 95% CI, 0.68-0.90). CONCLUSIONS AND RELEVANCEWe identified documented reasons for ACEI/ARB non-use (including both failures to initiate therapy and discontinuation after prior use) from unstructured text using an NLP method that does not require massive, expensive computing at inference time. By augmenting training data with schema-guided synthetic notes, we achieved robust, privacy-preserving performance within an NLP framework. This approach may support scalable clinical decision support systems to promote guideline-concordant prescribing.

4
Predicting Deprescribing of High-Risk Medications Using Provider EHR Use and Patient Characteristics

Gu, B.; Jungo, K. T.; Lauffenburger, J.; Choudhry, N.; Isaac, T.; Zambrano, J.; Yang, J.

2026-07-29 health informatics 10.64898/2026.07.25.26358746 medRxiv
Top 0.1%
59.4%
Show abstract

Potentially inappropriate medications expose older adults to preventable harm, yet deprescribing remains difficult to implement consistently. Although electronic health record (EHR) interventions can reduce prescribing, health systems lack clear evidence about which routinely captured patient, primary care provider (PCP), and intervention-design factors predict medication discontinuation or dose tapering. Understanding these determinants is essential for targeting and scaling deprescribing support. In this study, we conducted the first machine-learning analysis of these trial data. We analyzed 2,979 adults aged 65 years or older and 158 structured EHR features spanning patient characteristics, PCP characteristics and EHR-use behaviors, and deprescribing-tool design. We compared eight models for predicting medication discontinuation or dose tapering and used SHAP to examine feature importance. TabPFN achieved the highest positive predictive value (67.71%), AUROC (74.30%), and AUPRC (60.95%), although overall predictability was moderate. Our findings show that even rich structured EHR data only moderately predict deprescribing, suggesting that important clinical determinants are not captured in routine fields. PCP EHR-use measures accounted for 19 of the 25 highest-ranked TabPFN features, although they also constituted most candidate predictors. The study provides the health system with an informative reference for predicting high-risk medication deprescribing. Future models should incorporate richer clinical context and undergo external validation before informing personalized deprescribing support.

5
Current Limitations of Electronic Health Record Systems in Supporting Pragmatic Clinical Trials: Insights from the eMERGE Consortium

Wagholikar, K. B.; Pacheco, J. A.; Gordon, A. S.; Khan, A.; Khales, B. N.; Benoit, B.; Kerman, B. J.; Weng, C.; Ta, C.; Prows, C. A.; Johnson, R.; Roden, D. M.; Crosslin, D.; McNally, E. M.; Karlson, E. W.; Mentch, F.; Jarvik, G. P.; Wiesner, G. L.; Hakonarson, H.; Cimino, J. J.; Thayer, J. G.; Smoller, J. W.; Linder, J. E.; Connolly, J.; Peterson, J. F.; Cortopassi, J.; Kiryluk, K.; Hamed, M.; Maradik, M.; Puckelwartz, M. J.; Naderian, M.; Walton, N.; Limdi, N.; Maripuri, D. P.; Walunas, T.; Gainer, V.; Luo, Y.; Liu, C.; Kenny, E. E.; Espinoza, A.; Rowley, R.; Wei, W.-Q.; Murphy, S.

2025-04-03 health informatics 10.1101/2025.04.01.25325049 medRxiv
Top 0.1%
56.1%
Show abstract

Pragmatic clinical trials (PCTs) evaluate interventions in real-world settings, often using electronic health records (EHRs) for efficient data collection. We report on the challenges in performing EHR analysis of health-care provider orders in a PCT within the eMERGE consortium, which investigates the impact of reporting genome-informed risk assessments (GIRA) to over 25,000 patients across 10 academic medical centers. Clinical informaticians conducted a landscape analysis to identify approaches for evaluating the outcomes of GIRA reporting through the EHR. Of 98 identified outcomes, 54 (55.1%) were determined to be difficult to extract because they involved provider orders, which are typically documented in free text or proprietary formats within the EHR and only mapped to standardized codes after the service is completed. These findings highlight a critical barrier in using EHRs to support PCTs. The authors recommend closer collaboration between clinicians and informaticians, improved EHR systems that support standardized order entry, and future use of machine learning to automate analysis of provider behavior in clinical trials.

6
Augmenting Structured Diagnoses through Effective Use of Pre-trained Large Language Models on Clinical Notes

Razzaghi, H.; Nguyen, N.; Pargi, M.; Wieand, K.; Bunnell, T.; Bailey, C.

2026-06-02 health informatics 10.64898/2026.05.30.26354533 medRxiv
Top 0.1%
55.0%
Show abstract

Objective Clinical narrative provides a unique window into provider reasoning and attribution, but use has been limited by resource requirements and extensive fine-tuning, and LLMs in particular have traditionally not performed well at medical coding. We optimize and evaluate a reproducible method for automated diagnosis assignment using LLMs in clinical notes and compare with EHR structured diagnoses. Methods We used GPT-OSS for prompt engineering and task segmentation to create a model that extracts ICD-10-CM diagnoses, with estimates of severity, currency, and importance, from progress notes. We assessed performance across multiple cohorts of patients aged 0-21 years. For each, 100 outpatient provider notes were selected across levels of severity, along with coded diagnoses from that visit (EHR); a subset of 130 notes were subjected to clinical expert review. Results Comparison showed 18.7% exact code and 33.3% ICD-10-CM category match between EHR and LLM, but semantic similarity of 0.93 at the category level. Compared to expert review, LLM precision was 0.84 and recall 0.49 for exact matches, and 0.92 and 0.62, respectively, for category-level matching. In contrast, EHR coded diagnoses showed slightly higher precision (0.94 for both cases) and substantially lower recall (0.27 and 0.43) versus expert review. Codes not identified by the LLM were more often rated by the reviewer as lower importance or certainty. Conclusion We demonstrate a reusable approach to optimizing a pretrained LLM for use in diagnosis extraction from clinical notes, facilitating large-scale diagnosis screening by LLMs without the need for expensive study-specific model refinement.

7
Investigating Primary Care Indications to Improve the Quality of Electronic Health Record Data in Target Trial Emulation for Dementia

Sunog, M.; Magdamo, C.; Charpignon, M.-L.; Albers, M. W.

2025-04-10 neurology 10.1101/2025.04.08.25325485 medRxiv
Top 0.1%
53.6%
Show abstract

Missing data, inaccuracies in medication lists, and recording delays in electronic health records (EHR) are major limitations for target trial emulation (TTE), which uses EHR data to retrospectively emulate a clinical trial. EHR-based TTE relies on recorded data that proxy actual drug exposures and outcomes. While prior work has proposed various methods to improve EHR data quality, here we investigate the underutilized consideration that encounters with a primary care provider (PCP) may result in more accurate data in the EHR. Patients with a PCP within the EHR network being studied tend to have more encounters overall and a greater proportion of the types of encounters that yield comprehensive and up-to-date records. By contrasting data for patients with and without a PCP in the considered EHR network, we demonstrate how PCP status affects EHR data quality. Through a case study, we then empirically examine the impact on TTE of including a PCP status feature either in the propensity score and outcome models or as an eligibility criterion for cohort selection, versus ignoring it. Specifically, we compare the estimated effects of two first-line antidiabetic drug classes on the onset of Alzheimers Disease and Related Dementias. We find that the estimated treatment effect is sensitive to the consideration of PCP status, particularly when used as an eligibility criterion. Our work suggests that further researching the role of PCP status may improve the design of pragmatic trials. Data and Code AvailabilityThe study uses EHR data from the Research Patient Data Registry (Nalichowski et al., 2007), social vulnerability index (SVI) data from the Agency for Toxic Substances and Disease Registry (https://www.atsdr.cdc.gov/placeandhealth/svi), and Massachusetts death records from the Registry of Vital Records and Statistics. Because the data contain patient information, they cannot be made available. Institutional Review Board (IRB)This research was performed under MGB IRB approval (protocol 2023P000604).

8
Verifiable Summarization of Electronic Health Records Using Large Language Models to Support Chart Review

Verma, R.; Alsentzer, E.; Strasser, Z.; Chang, L.; Roman, K.; Gershanik, E.; Hernandez, C.; Linares, M.; Rodriguez, J.; Thakral, D.; Unlu, O.; You, J.; Zhou, L.; Bates, D.

2025-06-03 health informatics 10.1101/2025.06.02.25328807 medRxiv
Top 0.1%
53.4%
Show abstract

Information overload in electronic health records (EHRs) hampers clinicians ability to efficiently extract and synthesize critical information from a patients longitudinal health record, leading to increased cognitive burden and delays in care. This study explores the potential of large language models (LLMs) to address this challenge by generating problem-based admission summaries for patients admitted with heart failure, a leading cause of hospitalization worldwide. We developed an extract-then-abstract approach guided by disease-specific "summary bundles" to generate summaries of longitudinal clinical notes that prioritize clinically relevant information. Through a mixed-methods evaluation using real-world clinical notes, we compared physicians ability to answer patient-specific clinical questions with the LLM-generated summaries versus standard chart review. While summary access did not significantly reduce overall questionnaire completion time, frequent summary use significantly contributed to faster questionnaire completion (p = 0.002). Individual physicians varied in how effectively they leveraged the summaries. Importantly, summary use maintained accuracy in answering clinical questions (88.0% with summaries vs. 86.4% without). All physicians indicated they were "likely" or "very likely" to use the summaries in clinical practice, and 87.5% reported that the summaries would save them time. Preferences for summary format varied, highlighting the need for customizable summaries aligned with individual clinician workflows. This study provides one of the first extrinsic evaluations of LLMs for longitudinal summarization, demonstrating their potential to enhance clinician efficiency, alleviate workload, and support informed decision-making in time-sensitive care environments.

9
Can Electronic care planning using AI Summarization Yield equal Documentation Quality? (EASY eDocQ)

Dorr, D. A.; Weiskopf, N. G.; Sabile, J. M.; Young, E.; Mathis, D.; Weller, S.; Alex, J.; Doerr, K.; Bedrick, S.

2025-01-24 health informatics 10.1101/2025.01.22.25320982 medRxiv
Top 0.1%
52.1%
Show abstract

ImportanceData, information and knowledge in health care has expanded exponentially over the last 50 years, leading to significant challenges with information overload and complex, fragmented care plans. Generative AI has the potential to facilitate summarization and integration of knowledge and wisdom to through rapid integration of data and information to enable efficient care planning. ObjectiveOur objective was to understand the value of AI generated summarization through short synopses at the care transition from hospital to first outpatient visit. DesignUsing a de-identified data set of recently hospitalized patients with multiple chronic illnesses, we used the data-information-knowledge-wisdom framework to train clinicians and an open-source generative AI Large Language Model system to produce summarized patient assessments after hospitalizations. Both sets of synopses were judged blinded in random order by clinician judges. ParticipantsDe-identified patients had multiple chronic conditions and a recent hospitalization. Raters were physicians at various levels of training. Main outcomeAccuracy, succinctness, synthesis and usefulness of synopses using a standardized scale with scores > 80% indicating success. ResultsAI and clinicians summarized 80 patients with 10% overlap. In blinded trials, AI synopses were rated as useful 75% of the time versus 76% for human generated ones. AI had lower succinctness ratings for the Data synopsis task (55-67%) versus human (84-86%). For accuracy and synthesis, AI had near equal or better scores in other domains (AI: 72%-79%, humans: 68%-84%), with best scores from AI in Wisdom. Interrater agreement was moderate, indicating different preferences for synopsis content, and did not vary between AI and human-created synopses. DiscussionAI-created synopses that were nearly equivalent to human-created ones; they were slightly longer and did not always synthesize individual data elements compared to humans. Given their rapid introduction into clinical care, our framework and protocol for evaluation of these tools provides strong benchmarking capabilities for developers and implementers. Key PointsO_ST_ABSQuestionC_ST_ABSCan a Generative AI Large Language Model be trained to generate accurate and useful patient synopses through chart summarization for use in outpatient settings after hospital discharge? FindingsUsing a Data-Information-Knowledge-Wisdom framework, clinicians and an open-source AI system were trained to summarize charts; these synopses were rated blindly using a standardized index. Synopses from the AI were rated as useful 75% of the time versus 76% for human generated ones, AI synopses scored highest in Wisdom for accuracy and synthesis. Interrater agreement was moderate but did not vary between AI and human. MeaningThis study provides a concrete, replicable protocol for benchmarking LLM summarization outputs and demonstrates general equivalence to human-created synopses for outpatient use after care transitions.

10
The EHR Density Index: A new method to control for EHR data inconsistency across patients

Bhatia, A.; Lash, S.; McIntee, T.; Pfaff, E.

2026-08-06 health informatics 10.64898/2026.08.03.26359595 medRxiv
Top 0.1%
51.9%
Show abstract

Electronic health record (EHR) data vary substantially in documentation density across patients, independent of disease burden. Existing tools such as the Charlson Comorbidity Index (CCI) and Elixhauser Comorbidity Index measure disease burden but do not capture differences in data volume, leaving a common source of bias unaddressed in EHR-based analyses. To address this gap, we developed the EHR Density Index (EDI), which characterizes the quantity, depth, and breadth of EHR data per patient per year, normalized by utilization patterns, using records from 24,987 adult patients at UNC Health (2018 - 2024). The EDI combines a utilization cluster assigned via Gaussian Mixture Model with within-cluster residuals quantifying documentation volume across four clinical domains. Four interpretable clusters emerged; while CCI predicted cluster membership, its associations with within-cluster residuals were weak, confirming the EDI captures dimensions of the patient record distinct from disease burden. The EDI is intended as a covariate to address documentation density as a source of confounding in real-world data-driven research.

11
Implementation of Human-in-the-Loop ChatGPT-based Patient Screening Across Multiple Diverse Clinical Trials

Dohopolski, M.; Esselink, K.; Desai, N.; Grones, B.; Patel, T.; Jiang, S.; Peterson, E.; Navar, A. M.

2026-03-27 health informatics 10.64898/2026.03.20.26348890 medRxiv
Top 0.1%
51.2%
Show abstract

Purpose: Manual screening for trial eligibility is inefficient and costly. We prospectively evaluated a large language model (LLM)-assisted prescreening workflow across multiple active trials. Methods: We deployed a retrieval-augmented generation LLM-based pipeline across multiple trials at an academic medical center. Structured electronic health record data and free-text notes were used by the LLM to classify each criterion as either met, likely met, likely not met, not met, uncertain, or no documentation found, with accompanying rationale. Coordinators were provided a sorted patient list based on LLM-derived eligibility and reviewed each case, documenting their assessment of individual criteria and final prescreening status (success vs failure). Criterion-level performance--accuracy, sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), and F1 score--was calculated and tracked over time. Patient prescreening status was also evaluated as a function of the percentage of individual AI criteria met (60--80% and [&ge;]80%). Results: From October 2024--September 2025, 39,182 patients were prescreened using the LLM workflow across 26 studies (21 oncology and 5 non-oncology), encompassing 112 distinct criteria. A total of 914 patients with high likelihood of eligibility underwent coordinator review (5,096 criteria evaluated). Aggregated criterion-level performance was as follows: accuracy 0.94 (95% CI, 0.92--0.96), sensitivity 0.98 (0.97--0.99), specificity 0.81 (0.71--0.88), PPV 0.95 (0.92--0.97), NPV 0.93 (0.90--0.95), and F1 score 0.97 (0.95--0.97). Twenty-seven criteria prompts across 14/26 trials were automatically updated based on coordinator feedback. Patients with [&ge;]80% of AI-labeled criteria classified as met or likely met were more likely to be reviewed by coordinators (544/987, 55.1% vs 372/397, 93.7%) and more likely to be labeled as prescreening successes (104/544, 19.1% vs 162/372, 43.5%) compared to those with 60--80%. The average cost was $0.12 per patient. Conclusion: An LLM-assisted, human-in-the-loop prescreening workflow demonstrated high criterion-level performance at low cost across a diverse set of actively enrolling clinical trials. Structured coordinator feedback enabled an automated learning system, improving screening efficiency while preserving necessary human oversight.

12
Sharing Aggregated Patient Counts in Place of Line-Level EHR Data: Analytic Fidelity and the Limits of Count Suppression for Privacy

Chen, Y.; McMurry, A.; Gottlieb, D.; Jones, J. R.; Strober, B. J.; Mandl, K. D.

2026-08-21 health informatics 10.64898/2026.08.18.26359984 medRxiv
Top 0.1%
51.2%
Show abstract

Objective. Privacy regulation constrains sharing line-level electronic health records (EHR) across institutions. One alternative is to aggregate counts into a cube, a table of counts for every combination of categorical variables, with cells below a threshold suppressed. This study asked whether common analyses on the cube reproduce conclusions from line-level data, and whether suppression prevents recovery of the small cells it is meant to hide. Materials and Methods. A Bayesian count-inference pipeline was built that reconstructs suppressed counts and doubles as a reconstruction attack. Applied to 285 pediatric kidney-transplant patients at Boston Children's Hospital, statistical fidelity (Jensen-Shannon divergence, Cramer's V, and R2) and analytical utility (marginal distributions, subgroup graft rejection odds ratios, and logistic-regression classification) were evaluated. Conditional Tabular GAN (CTGAN) synthetic data served as a comparator. Results. Statistical analyses on the cube recapitulated results from line-level data. Across 106 demographic-by-medication subgroups, a bootstrap mean of 3.5 subgroups showed a significant graft-rejection association. The cube's odds-ratio sign changes reversed no significant associations, versus 2.3 for CTGAN. The same reconstruction also defeated suppression: in a 10-variable cube, 76.6% of suppressed cube cells were recovered exactly (14,554 of 18,994), including 85.5% of single-patient cells. Discussion. The cube reproduced common kidney-transplant analyses, but the same reconstruction also recovered suppressed cells; fidelity and privacy risk are thus two faces of one reconstruction rather than independent properties. Conclusions. The cube is a useful surrogate for these kidney-transplant analyses only when paired with a stronger privacy mechanism. This study demonstrated reconstructability of suppressed counts, not re-identification.

13
Large Language Model Augmented Clinical Trial Screening

Beattie, J.; Owens, D.; Navar, A. M.; Schmitt, L. G.; Taing, K.; Neufeld, S.; Yang, D.; Chukwuma, C.; Gul, A.; Lee, D. S.; Desai, N.; Moon, D.; Wang, J.; Jiang, S.; Dohopolski, M.

2024-08-28 health informatics 10.1101/2024.08.27.24312646 medRxiv
Top 0.1%
50.0%
Show abstract

PurposeIdentifying potential participants for clinical trials using traditional manual screening methods is time-consuming and expensive. Structured data in electronic health records (EHR) are often insufficient to capture trial inclusion and exclusion criteria adequately. Large language models (LLMs) offer the potential for improved participant screening by searching text notes in the EHR, but optimal deployment strategies remain unclear. MethodsWe evaluated the performance of GPT-3.5 and GPT-4 in screening a cohort of 74 patients (35 eligible, 39 ineligible) using EHR data, including progress notes, pathology reports, and imaging reports, for a phase 2 clinical trial in patients with head and neck cancer. Fourteen trial criteria were evaluated, including stage, histology, prior treatments, underlying conditions, functional status, etc. Manually annotated data served as the ground truth. We tested three prompting approaches (Structured Output (SO), Chain of Thought (CoT), and Self-Discover (SD)). SO and CoT were further tested using expert and LLM guidance (EG and LLM-G, respectively). Prompts were developed and refined using 10 patients from each cohort and then assessed on the remaining 54 patients. Each approach was assessed for accuracy, sensitivity, specificity, and micro F1 score. We explored two eligibility predictions: strict eligibility required meeting all criteria, while proportional eligibility used the proportion of criteria met. Screening time and cost were measured, and a failure analysis identified common misclassification issues. ResultsFifty-four patients were evaluated (25 enrolled, 29 not enrolled). At the criterion level, GPT-3.5 showed a median accuracy of 0.761 (range: 0.554-0.910), with the Structured Out-put + EG approach performing best. GPT-4 demonstrated a median accuracy of 0.838 (range: 0.758-0.886), with the Self-Discover approach achieving the highest Youden Index of 0.729. For strict patient-level eligibility, GPT-3.5s Structured Output + EG approach reached an accuracy of 0.611, while GPT-4s CoT + EG achieved 0.65. Proportional eligibility performed better over-all, with GPT-4s CoT + LLM-G approach having the highest AUC (0.82) and Youden Index (0.60). Screening times ranged from 1.4 to 3 minutes per patient for GPT-3.5 and 7.9 to 12.4 minutes for GPT-4, with costs of $0.02-$0.03 for GPT-3.5 and $0.15-$0.27 for GPT-4. ConclusionLLMs can be used to identify specific clinical trial criteria but had difficulties identifying patients who met all criteria. Instead, using the proportion of criteria met to flag candidates for manual review maybe a more practical approach. LLM performance varies by prompt, with GPT-4 generally outperforming GPT-3.5, but at higher costs and longer processing times. LLMs should complement, not replace, manual chart reviews for matching patients to clinical trials.

14
From Conversation to Chart: An Analysis of Clinician Edits to Ambient AI Draft Notes

Guo, Y.; Hu, D.; Zhou, Y.; Lyu, T.; Sutari, S.; Tam, S.; Chow, E.; Perret, D.; Pandita, D.; Zheng, K.

2026-01-06 health informatics 10.64898/2026.01.05.26343471 medRxiv
Top 0.1%
49.0%
Show abstract

Structured AbstractO_ST_ABSObjectiveC_ST_ABSAmbient artificial intelligence (AI) tools are increasingly adopted in clinical practices. This study investigated whether and how clinicians edit AI-generated drafts and the linguistic differences between AI drafts and clinician-finalized notes. Materials and MethodsThis retrospective study analyzed real-world data from ambulatory clinics at a large academic health system spanning two vendor deployments. We quantified clinicians editing behavior using the Myers diff algorithm to compare AI drafts and final documentation. We then applied statistical and linguistic analysis to study factors associated with the frequency/intensity of editing across note sections, turnaround time, clinician characteristics, and encounter types. ResultsAcross 23,760 notes that included one or more ambient AI sections, 84.4% were edited by clinicians before signing off. While rates of unedited notes differed across note sections and care settings, the dominant source of variation was individual clinician practice style rather than specialty-level norms. Notes signed after 24 hours had lower overall edit intensity. The final versions showed small but statistically significant linguistic changes and exhibited slightly higher lexical diversity and modest changes in readability. Editing is most intensive in the assessment and plan section, and varies across specialties. Conclusion and DiscussionA majority of AI-drafted clinical notes were edited by clinicians, although the editing rate varies across note sections, medical specialties, and individual clinicians. Future research is needed to further analyze this editing behavior to inform improvement in AI-assisted clinical documentation to achieve better documentation quality, efficiency, and clinician satisfaction.

15
Randomized Trial Protocol: Epic Generative AI Chart Summarization Tool to Reduce Ambulatory Provider Cognitive Task Load

Chin, A. T.; Zhu, N.; Kingsley, T. C.; Mynampati, P.; Phipps, Y.; Romanov, A.; Vangala, S.; Weng, M.; Wisk, L. E.; Woo, H.; Mafi, J. N.; Lukac, P. J.

2026-02-22 health informatics 10.64898/2026.02.20.26346503 medRxiv
Top 0.1%
48.9%
Show abstract

BackgroundEHR documentation and chart review contribute to clinician workload and burnout. To alleviate pre-charting burden, Epic has released a new generative AI chart summarizer tool, which has become widely adopted; however, its impact has not been examined in randomized trials. ObjectiveTo evaluate whether access to an Epic generative AI chart summarization tool reduces cognitive task load among ambulatory providers compared with usual care. MethodsTwo-arm, parallel-group randomized controlled trial among ambulatory clinicians across multiple specialties. Clinicians will be randomized 1:1 to tool access versus usual care for 90 days. The primary outcome is change in a 4-item physician task load (PTL) adapted for the pre-charting task. Exploratory outcomes include EHR-derived time metrics (Caboodle and Signal), professional fulfillment/burnout (PFI), usability (SUS), clinician satisfaction, aggregated patient experience item from CG-CAHPS, and reported safety related metrics. Ethics and DisseminationAnalyses will use clinician-level survey responses and aggregated EHR metrics; no patient-level protected health information will be included in the analytic dataset. Results will be disseminated via preprint and peer-reviewed publication. Article summary - Strengths and limitations of this studyO_LIThis study is a 3-month pragmatic randomized controlled trial evaluating a native EHR-embedded generative AI tool that summarizes prior clinical notes for ambulatory encounters. C_LIO_LIThe primary outcome uses a validated cognitive task load instrument adapted specifically for pre-charting activities. C_LIO_LIExploratory outcomes include objective EHR-derived time metrics, validated psychometric measures of burnout and professional fulfillment, and clinician-reported survey measures assessing perceived usefulness of the tool. C_LIO_LIThe trial is single-centered, which may limit generalizabilty, and the intervention is optional-use and unblinded, which may attenuate observed effects and introduce performance bias. C_LI

16
Development and validation of an algorithm to identify front-line clinicians using EHR audit log data

Baratta, L. R.; Wang, J.; Osweiler, B. W.; Lew, D.; Eiden, E.; Kannampallil, T. G.; Lou, S. S.

2026-02-16 health informatics 10.64898/2026.02.13.26346268 medRxiv
Top 0.1%
48.7%
Show abstract

BackgroundInterprofessional teams are central to high quality patient care. However, identifying the clinician primarily responsible for a patient requires labor-intensive methodologies. Although electronic health record (EHR) audit logs offer a scalable alternative, its use for identifying frontline clinicians is underdeveloped. ObjectiveTo develop and validate an algorithm utilizing EHR audit logs to identify the primary frontline clinician per patient day of an encounter and to describe care continuity patterns. MethodThis was a cross-sectional cohort study of adult inpatient medicine encounters at 12 hospitals in a single health system using a shared EHR. Admissions from February 1, 2023-April 30, 2023, with length of stay of at least 3 days and without an intensive care unit admission were included. Four algorithm iterations were designed to identify the attending physician, resident, or advanced practice provider primarily responsible for patient care on each patient-day. Performance of each algorithm was compared with manual chart review on 1,401 patient-days from 246 randomly sampled patient encounters. Accuracy between an algorithm and the chart review standard was compared using McNemars test with Bonferroni adjusted p-values. ResultsThe best performing algorithm correctly identified the primary clinician responsible for patient care on 91% of patient-days (1,268/1,401), outperforming the naive approach using frequency of actions (78% accuracy, 1,098/1,401, p<0.001). Algorithm errors were attributable to misidentified specialty and ambiguity on days with transitions of care or shared responsibilities between clinicians. The best performing algorithm was applied to the entire cohort (5,801 encounters and 34,001 patient-days) where it identified attending physicians, resident physicians, and APPs as the frontline clinician for 26,750 (79%), 3,106 (9%), and 4,145 (12%) of patient days respectively. Each encounter had a median of 1 (IQR 0-2) handoff between frontline clinicians. ConclusionsWe developed a scalable, audit log-based algorithm to determine the front-line clinician with excellent accuracy compared with manual chart review.

17
MedSDoH: A Rule-Based System for Extracting Social Determinants of Health from Multi-site EHRs Based on the OHNLP Framework

Ahn, J.; Fu, S.; Palacios, D. M.; Jeong, H.-H.; Wang, L.; Swartz, M. C.; Tosur, M.; Redondo, M. J.; Wu, X.; Yue, Z.; Kakadiaris, A.; Wang, N.; Li, Z.; Huang, M.; Wen, A.; Harris, D.; Wang, Y.; Kwak, M. J.; Liu, Z.; Liu, H.

2026-04-29 health informatics 10.64898/2026.04.27.26351699 medRxiv
Top 0.1%
46.4%
Show abstract

ObjectiveSocial Determinants of Health (SDoH) are critical to patient care and population health. Despite their importance, SDoH information is frequently embedded within unstructured clinical text such as patient-reported information or social worker notes, which limits its use on clinical decision-making and resource allocation. Although transformer-based models represent the current state of the art, their scalability, computational requirements, and limited transparency pose barriers to large-scale multi-site clinical implementation. In this context, rule-based NLP systems remain valuable, particularly when explainability, reproducibility, and rapid customization are essential. MethodsMedSDoH was developed within the Open Health Natural Language Processing (OHNLP) Framework using literature-derived SDoH resources, standardized domain definitions, and expert-curated rulesets. Large language models (LLMs) were used during development to assist with rule generation and lexicon expansion. Rules were iteratively refined against a gold-standard annotated corpus from two health systems and then evaluated on independent datasets. ResultThe final system included 942 regular expression rules spanning 22 SDoH domains. On validation on two external datasets, MedSDoH demonstrated generalizability and comparable performance across sites. The system has been made publicly available so research community can collaboratively contribute to the maintenance and extension through disease- or site-specific adaptations. ConclusionMedSDoH is a computationally efficient and open-source system for large-scale SDoH extraction from clinical text. It is well-suited for multi-site adaptation and deployment in resource-constrained settings.

18
Evaluating an LLM-Assisted Workflow for Clinical Documentation: A Pilot Randomized Controlled Trial on Time and Quality

Takayama, T.; Sado, K.; Suda, K.; Tamura, H.; Ueda-Arakawa, N.; Ishihara, K.; Ueda, Y.; Okamoto, K.; Santos, L. H. d. O.; Oshika, T.; Kuroda, T.; Tsujikawa, A.; Miyake, M.

2025-10-07 health informatics 10.1101/2025.10.06.25337211 medRxiv
Top 0.1%
46.2%
Show abstract

IMPORTANCELarge language models (LLMs) have been investigated for clinical documentation, with concerns about hallucinations and factual errors. Clinician review and revision of LLM-generated drafts are therefore considered essential, yet the impact of such workflow on both documentation time and quality remains unknown. OBJECTIVETo assess whether physician review and editing of LLM-generated drafts improves the time and quality of clinical documentation compared with clinician-only drafting in a randomized controlled trial setting. DESIGNSingle-center, parallel-group, prospective, randomized, open-label, blinded-endpoint (PROBE) pilot trial conducted from February 18 to March 14, 2025. SETTINGKyoto University Hospital, Department of Ophthalmology. PARTICIPANTSTwenty-one ophthalmology physicians were randomized; 17 completed the study, and 4 withdrew before initiating intervention. INTERVENTIONSParticipants were randomized to either the Clinician-in-the-loop group or the Clinician-only group. All participants created discharge summaries and referrals for six simulated patient records. In the Clinician-in-the-loop group, drafts were generated with an LLM assistant and then reviewed and edited by participants, whereas in the Clinician-only group, documents were drafted from scratch using matched templates. Unedited LLM drafts were additionally analyzed as the LLM-only group. MAIN OUTCOMES AND MEASURESDocument creation time (primary) and expert-rated document quality across six domains plus overall quality (secondary). RESULTSSeventeen physicians submitted 48 discharge summaries and 48 discharge referrals in the Clinician-in-the-loop group, 54 of each document type in the Clinician-only group, and 48 of each in the LLM-only group. For summaries, Clinician-in-the-loop was associated with shorter creation time versus Clinician-only ({beta} = -59.4 seconds; 95% CI, -118.1 to -0.8; P = .047). For referrals, clinician-in-the-loop required more time ({beta} = 94.8 seconds; 95% CI, 40.4 to 149.3; P<.001). In most quality domains for both document types, the clinician-in-the-loop workflow outperformed clinician-only drafting. LLM-only drafts were fastest but had the lowest quality. CONCLUSIONS AND RELEVANCEA clinician-in-the-loop approach improved document quality and accelerated documentation. Active clinician review of LLM-generated drafts is essential for clinical documentation, and such workflows may help enhance working conditions and patient care. TRIAL REGISTRATIONClinicalTrials.gov Identifier NCT07187050 Key PointsO_ST_ABSQuestionC_ST_ABSDoes a workflow in which clinicians review and edit LLM-generated drafts ("clinician-in-the-loop") improve the efficiency and quality of clinical documentation compared with clinician-only drafting? FindingsIn this single-center randomized controlled trial including 17 ophthalmology physicians, discharge summaries were significantly faster with the clinician-in-the-loop. Across both document types, clinician-in-the-loop drafts achieved higher quality scores than clinician-only documents. MeaningA clinician-in-the-loop workflow using LLMs can simultaneously enhance the efficiency and quality of clinical documentation; accordingly, active clinician review remains essential for improving working conditions and patient care overall.

19
A Quantitative Framework for EHR Cohort Refinement.

Songthangtham, N.; Simon, G.; Johnson, S. G.

2026-04-28 health informatics 10.64898/2026.04.27.26351837 medRxiv
Top 0.1%
45.5%
Show abstract

Electronic Health Record (EHR) based research depends on accurate cohort definitions, yet current workflows offer little early insight into cohort quality and rely heavily on slow, ad hoc manual review. We developed a quantitative, iterative framework that integrates model-guided case selection with sequential statistical testing to provide an early, data-driven signal of cohort accuracy. Using an OMOP-standardized dataset and the PCORnet Type 2 Diabetes phenotype as the gold-standard proxy, we evaluated eight sampling strategies across starting review sizes from 30 to 2,000 cases. At each iteration, the framework retrained an internal logistic regression model, selected cases for review, and applied a Bonferroni-adjusted Agresti-Coull upper bound to assess whether the cohort met a pre-specified accuracy threshold. Across 48 simulation conditions, starting sample size strongly shaped iteration count and total review burden, and adaptive sampling strategies consistently required fewer reviewed cases than fixed-batch methods. These findings demonstrate a reproducible, statistically grounded approach for refining EHR cohort definitions while reducing manual review effort.

20
What Do Clinicians Edit in Ambient AI-Drafted Clinical Documentation? A Qualitative Content Analysis

Guo, Y.; Hu, D.; Yang, Z.; Kim, S.; Tran, B.; Lee, J.; Vallabhaneni, S.; Zehrung, R.; Sutari, S.; Tam, S.; Chow, E.; Perret, D.; Pandita, D.; Zheng, K.

2026-01-06 health informatics 10.64898/2026.01.05.26343473 medRxiv
Top 0.1%
45.4%
Show abstract

ObjectiveAmbient artificial intelligence (AI) documentation is increasingly used to draft clinical notes from patient-provider conversations, but how clinicians revise and finalize these drafts is not well understood. This qualitative content analysis study characterizes real-world edits to AI-generated drafts and identifies opportunities for improvement of AI design and the implementation process. Materials and MethodsEight coders analyzed clinical documentation generated by ambient AI from 200 clinical encounters. We developed an inductive coding framework with 11 codes across three categories: clinical content, terminology, and language style. Interrater reliability was assessed using Cohens kappa. We then applied thematic analysis to synthesize patterns across the coded edits. ResultsThe most frequently edited content pertained to clinical facts including orders (e.g., procedures, lab tests) (40.0%), symptoms (30.3%), medication prescriptions (27.3%), and diagnosis descriptions (25.9%). In comparison, edits related to terminology use (11.6%) and language style (7.2%) were less frequent. The results of our thematic analysis show that most edits can be categorized into one of the following five types: to correct factual errors, to address needs of medical specialty, to express diagnostic certainties, to convert patient expressions into objective assessments recorded in medical terms, and to reorganize or condense content. Conclusion and DiscussionClinicians routinely revise ambient AI drafts to improve accuracy and clinical specificity. Future work on AI development and clinical implementation should emphasize specialty customization and support personalized documentation practices, alongside clinician education that promotes robust and consistent review routines to ensure documentation quality.